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® Algorithm for identifying bad data. 

@ An algorithm, and methodology for its application, useful in digital computer systems incorporating read- 
modify-write data storage systems to accurately identify rewritten data which has been determined bad before 
being rewritten. 
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BAD DATA ALGORrTHM 

Field of the Invention 



The present invention relates to bad data identification procedures (or digital computer systems, and 
5 more particularly to an algorithm for ensuring that known bad data rewritten and stored In memory can be 
identified as bad even if an additional single bit error occurs. 

Background of the invention 

10 

For read-modity-write data storage systems in digital computers, it is desirable to include a provisio n for 
preventing data discovered as bad and uncorrectab l e during the read process from being re written and later 
read_as_being-correctable_jdata- Such an occurrence is generally the result of the rewritten bad data 
J 5 inadvertently acquiring an additional bit error, which then make certain single bit errors occurring in the 
stored bad data look like correctable data when reread. The algorithms generally in use for this purpose do 
not protect the stored bad data from all cases of such single bit error influence. 



20 Summary of the Invention 

The present invention assures proper identification of stored data already determined to be bad. after 
such bad data Is rewritten as part of a read-modify-write operation. The complete data word comprises 40 

25 bits with 7 check bits for error correction combined with a special mark bit and 32 bits of data. After read 
data is determined to be uncorrectably bad, 7 check bits are rated according to the error correction code 
<ECC) employed, and then the check bits are inverted. A special mark bit is also added to the data bits and 
inverted check bits. The data bits, inverted check bits and mark bit are then rewritten. When the bad data is 
reread, a new set of 7 check bits are generated, and the new check bits are compared to the 7 inverted 

30 check bits in an exclusive OR relationship, which when correlated with the special mark bit, provides an 
accurate indication that the condition of the rewritten data is bad. even if an additional single bit error occurs 
during the rewrite or read process. 

35 Description of the Drawings 

Figure 1 is a typical digital computer system including a read-modlfy-write memory system suitable 
for incorporating the present invention. 
40 Figure 2 is a flow chart of the basic methodology of the present invention applied to the read-modify- 

write memory system used in the digital computer system shown in Rgure 1 . 

Figure 3 is a flow chart of the specific methodology for developing the algorithm used in the present 
invention shown in Figure 2. 

45 Detailed Description of the Preferred Embodiment 

Referring to the drawings, wherein like characters designate like or corresponding parts throughout the 
views. Figure 1 illustrates schematically a typical digital computer system 2 suitable for incorporating the 
50 present invention. The computer system 2 includes a central processing unit (CPU) 4, a memory system 6, 
an error con-ecting system (ECS) 8 and a buffer 1 0. System data is transmitted between the CPU 4 and the 
buffer 10 via a computer data bus 12. Address information is sent from the CPU 4 to the buffer 10 via an 
address bus 14. Likewise, address information is sent from the CPU 4 to the memory system 6 via an 
address bus 16. Communication between the buffer 10 and the memory system 6 Is provided by a 
communication data bus 18. Likewise, communication between the buffer 10 and the error correcting 
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system (ECS) 8 is provided by a communicafton data bus 20. Each of the above are well known 
components which may be Interconnected in a variety of well known configurations and are shown in block 
form for purposes of representation only, since they do not in themselves constitute part of the present 
invention. The memory system 6 may. for example, store 40 fait wide data words which include 32 data bit 

s positions. The buffer 10 may, for example, be a multiple data word buffer including a plurality of 40 bit 
addressable word buffers. Data Is transferred from the CPU 4 to the buffer tO and then to the memory 6. 
Likewise, data read from the memory 6 Is transferred to the buffer 10 and then to the CPU 4. 

When provisions are made for 40 bit data words each having 32 data bit positions. 8 bits are available 
for the error detection and correction process. The 8 bits may be employed in the ECS 8 to automatically 

w correct a single-bit error in any one of the data bit positions of each word read from the memory system 6. 
The ECS 8 may use any, or all. of the available bit positions for the error detection and correction process 
to allocate check bits for error correction purposes, the check bits being generated with the data of each 
data word according to any of the well known ECC codes. The check bits so allocated may then be 
employed to generate an ECC syndrome for error detection and correction using methods well known in the 

IS art. However, multiple errors detected In ttie data word during the ECC detection process, or data received 
from the CPU 4 for modifying ttie data word In which a parity error is discovered, cannot be corrected, and 
the data word must be designated as bad data when it is rewritten in the memory system 6. 

When data words are read from the memory system 6 as part of a read-modify-write process In 
response to command information from the CPU 4. the implementation of a bad data algorithm according to 

20 the present invention follows the flow chart shown in Figure 2. After the data word is read from the memory 
system 6 and new data is received with command information from the CPU 4 via the buffer 10, the data 
word from the memory system 6 and the new data from the CPU 4 are checked for errors. If no error is 
found in the data word from the memory 6 or in the new data from the CPU 4, then the read data is 
modified by the new data in accord with the command information from CPU 4, a new set of 7 check bits 

25 according to the ECC code used is generated and the mark bit Is set to the "0" slate. The mark bit position . 
is used to designate the status of the data, the mark bit being set to a "1 " state for bad data, and a "0" 
state for good data. The entire 40 bit data word including the 32 data bits modified with new data according 
to the command information from the CPU 4 the new 7 check bits and the "0" state mark bit is rewritten 
into the memory system 6. 

30 If error is found, but it is correctable, such as is the case if a single bit error is found In the read data 
word from the memory 6, the process continues the same as If no error was detected as described above, 
the modified data word being rewritten with 7 new check bits and a "0" state mark bit. If error Is found 
which Is uncorrectable, the data word from the memory 6 Is not modified by the new data from the CPU 4 
in accord with the command information, but seven new check bits are generated just as described above 

35 for data with no detected error, the check bits are then Inverted, and the mark bit is set to the "1" state. 
The 40 bit data word including the unmodified 32 data bits, the inverted check bits and a "1 " state mark bit 
Is then rewritten into the memory system 6. 

The data portion of data words designated bad data with the mark bit set in the "1 " state are rewritten 
as read, in non-Inverted form, because the detected error in such bad data may have been caused by a 

40 faulty dynamic random access memory {DRAM) chip in the memory system 6 that will always produce the 
same state. By writing back the same state to that faulty DRAM, rereading of the data as written is ensured. 
Otherwise, an additional single bit en-or might be read during the rereading process, which could cause the 
ECS 8 to erroneously Identify the data word as having a correctable error when the data word has actually 
already been identified as bad data. 

45 The method or process for detecting the bad data algorithm according to the present invention to 
correctly Identify rewritten data detected and designated as bad follows the flow chart shown in Figure 3. 
The data word read form the memory system 6 is stored in the ECS 8 and 7 new check bits are generated 
from the stored data according to the ECC code. The new check bits are then compared to the inverted 
check bits which fomn part of the stored data word in an exclusive OR relationship. The inverted stored 

so check bits produce an ECC syndrome which is a compliment of the syndrome that would normally be 
generated by a single bit error. The complimentary syndrome generated with the inverted stored check bits 
indicates uncorrectable multiple bit error more accurately than the normally generated syndrome. The 
complimentary syndrome will not accurately Identify data words with uncorrectable multifait error in one 
situation. Therefore, reliance cannot be placed only on the complimentary syndrome to determine that a 

5S read data word contains uncorrectable errors and the mark bit must be used. 

The one situation referred to in the preceding paragraph occurs when the initial bad data detected in 
the read data word from the memory system 6 contains one error in a data bit position and another in a 
check bit position and this bad data is rewritten into the memory system 6 with inverted check bits and the 



3 



EP 0 381 885 A2 



mark bit set to the "1" state as described above and then when this data is read an additional single bit 
error occurs in one of the bad data bit positions. This exception occurs because such bad data when reread 
can generate an ECC syndrome of 3. 4 or 5 "1"'s. A syndrome of 3 "1"'s indicates a single bit error. A 
single bit error is normally correctable. However, the correlation of the mark bit with the generated ECC 
5 syndrome in the decoding logic allows the reread data word to be correctly determined to have bad, 
uncorrectable data, since the mark bit is set to the "1" state for bad data and the "0" state for good data. 

The Table indicates all possibilities of error detected in read and reread data words initially read as 
having bad data. 

jo Table 
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(1) 3-1 "'s normally detected as a SEE but the mark bit distinguishes it as a MBE. 

(2) if the syndrome is 3 "1 "'s and the mark bit is 0 this will be reported incorrectly as a SBE. 



55 Thus, there has been described herein a bad data identification algorithm uniquely suited for read- 
modify-write memory storage systems in digital computers which can identify all single bit errors that occur 
on data marked bad and report them as uncorrectable multiple bit errors, while distinguishing single bit 
errors on good data. It will be understood that various changes in the details, arrangements and 
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configurations of \he parts and assennblies which have been described and illustrated above in order to 
explain the nature of the present invention may be made by those skilled in the art within the principle and 
scope of the present invention as expressed in the appended claims. 

s 

Claims 

1 . For a digital computer system including a memory system with a read-modify-write operation mode 
having a plurality of stored data words, each stored data word having a plurality of data bit positions and a 

w plurality of error correction code (ECC) bit positions, a method of identifying data words as bad data words 
which have uncorrectable errors, including the steps of: 

aliocating a bit of said ECC bit positions in each of said data words to indicate that said data has been 
determined uncorrectable; 

generating a mark bit for each of said data words in said allocated bit position; 
js generating check bits for each of said data words according to an error correction code (ECC) to fill the 
remainder of said ECC bit positions; 

inverting said generated check bits for each of said bad data words to form a plurality of corresponding 
inverted check bits; 

writing said data for each of said data words into its corresponding data bit positions; 
20 writing said inverted check bits for each of said bad data words into said corresponding ECC bit positions; 
writing said mark bit for each said data word into said allocated bit position; 
reading each of said written data words; 

generating new check bits for each of said read data words according to said ECC; 

comparing said read data generated check bits with said corresponding written check bits in an exclusive 
25 OR relationship to develop an ECC syndrome; and 

correlating said ECC syndrome with said mark bit in each of said read data words to identify that each of 
said bad data words contain uncorrectable errors. 

2. The method recited in claim 1. wherein said step of generating said mark bit includes generating said 
mark bit as a logical "1 " state for each said data word including said bad data. 

30 3. The method recited in claim 2. vyherein said data words comprise 40 bit data words. 

4. The method recited in claim 2. wherein said data bit positions of each said data word comprise 32 
data bit positions. 

5, The method recited in claim 4, wherein said step of generating check bits includes generating 7 
check bits for each of said data words. 

35 6. The method recited in claim 2, wherein said data words include good data words with good data and 
said step of generating said mark bit further includes generating said mark bit as a logical "0" state for 
each said good data word having good data. 

7. The method recited in claim 6 further including the steps of: 

writing said good data for each of said good data words into its corresponding data bit positions; and 
10 writing said check bits for each of said good data words into said corresponding ECC bit positions. 

8. The method recited in claim 7, wherein said data words comprise 40 bit data words. 

9. The method recited in claim 8. wherein said data bit positions of each data word comprise 32 data bit 
positions. 

10. The method recited in claim 9, wherein said step of generating check bits includes generating 7 
45 check bits for each of said data words. 

11. For a digital computer system including a memory system with a read-modify-write operation mode 
having a plurality of good and bad stored data words, each stored data word having a plurality of data bit 
positions and a plurality of error correction code {ECC) bit positions, a method of identifying data words as 
bad data words which have uncorrectable errors, including the steps of: 

50 allocating a bit of said ECC bit positions in each of said data words to indicate that said data has been 
determined uncorrectable; 

generating a logical "0" mark bit for each of said good data words in said allocated bit position: 
generating a logical "1" mark bit for each of said bad data words in said allocated bit position; 
generating check bits for each of said data words according to an error correction code (ECC) to fill the 
5S remainder of said ECC bit positions: 

inverting said generated check bits for each of said bad data words to fomn a plurality of corresponding 
inverted check bits; 

writing said data for each of said data words into its corresponding data bit positions; 
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writing said check bits for each of said good data words into said corresponding ECC positions; 
writing said inverted check bits for each of said bad data words into said corresponding ECC bit positions; 
writing said mark bit for each said data word into said allocated bit position; 
reading each of said written data words; 
5 generating new check bits for each of said read data words according to said ECC; 

comparing said read data generated check bits with said corresponding written check bits in an exclusive 
OR relationship to develop an ECC syndrome: and 

correlating said ECC syndrome with said mark bit in each of said read data words to identify that each of 
said bad data words contain uncorrectable en-ors. 
10 12. The method recited in claim 11 . wherein said data words comprise 40 bit data words. 

13. The method recited in claim 12, wherein said data bit positions of each said data word comprise 32 
data bit positions. 

14. The method recited in claim 13. wherein said step of generating check bits includes generating 7 
check bits for each of said data words. 
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